Papers with Knowledge Representation

198 papers
Improving Medical NLI Using Context-Aware Domain Knowledge (2020.starsem-1)

Copied to clipboard

Challenge: Domain knowledge is important to understand both the lexical and relational associations of words in natural language text . lack of annotated dataset can lead to wrong inference predictions .
Approach: They propose a knowledge adaptive approach that encodes the premise/hypothesis texts by leveraging supplementary external knowledge alongside the UMLS based on the word contexts.
Outcome: The proposed model can align token-level interactions between the premise and hypothesis more effectively.
The INCEpTION Platform: Machine-Assisted and Knowledge-Oriented Interactive Annotation (C18-2)

Copied to clipboard

Challenge: INCEpTION is an annotation platform for interactive and semantic annotation . the platform is both generic and modular .
Approach: INCEpTION is an annotation platform for interactive and semantic annotation . the platform incorporates machine learning capabilities which actively assist annotators .
Outcome: INCEpTION is an open-source annotation platform for tasks including interactive and semantic annotation.
Automatic Construction of Enterprise Knowledge Base (2021.emnlp-demo)

Copied to clipboard

Challenge: Existing knowledge bases are often based on bootstrapping entities from human-curated sources such as Wikipedia.
Approach: They propose to build a knowledge base from enterprise documents with minimal human intervention by using deep learning models and classical machine learning techniques.
Outcome: The proposed system is currently serving as part of a Microsoft 365 service.
Enabling Search and Collaborative Assembly of Causal Interactions Extracted from Multilingual and Multi-domain Free Text (N19-4)

Copied to clipboard

Challenge: a new searchable knowledge graph allows users to search for causal interactions in multiple languages . a recent study shows that search tools are shallow and do not support multilingual research .
Approach: They propose a system that integrates causal interactions into a single searchable knowledge graph.
Outcome: The proposed system extracts over 600 thousand causal statements from 120 thousand Portuguese publications with a precision of 62%.
Scalable Construction and Reasoning of Massive Knowledge Bases (N18-6)

Copied to clipboard

Challenge: Existing knowledge mining systems assume abundant human annotations for training high quality machine learning models, which is impractical when trying to deploy IE systems to a broad range of domains, settings and languages.
Approach: They introduce how to extract structured facts from text corpora to construct knowledge bases.
Outcome: The proposed methods are weakly-supervised and domain-independent for knowledge base construction across various domains.
Scalable graph-based method for individual named entity identification (D19-53)

Copied to clipboard

Challenge: Named entity recognition (NED) is a method for identifying named entities within a knowledge base.
Approach: They propose a method for individual identification requiring few annotated data samples.
Outcome: The proposed method is well-motivated for integration in real systems.
CL Scholar: The ACL Anthology Knowledge Graph Miner (N18-5)

Copied to clipboard

Challenge: ACL Anthology is a repository for papers related to computational linguistics and natural language processing.
Approach: They propose to automate periodic crawling, indexing and processing of new articles . they propose to use CL Scholar to support more than 1200 natural language queries .
Outcome: The proposed system can answer three different types of natural language queries.
Automatic Taxonomy Induction and Expansion (D19-3)

Copied to clipboard

Challenge: Knowledge Graph Induction Service (KGIS) enables automatic taxonomy induction and human-in-the-loop curation.
Approach: They describe the features of the Knowledge Graph Induction Service (KGIS) KGIS allows the user to semi-automatically curate and expand the induced taxonomies through a component called Smart SpreadSheet .
Outcome: The Knowledge Graph Induction Service (KGIS) is an end-to-end knowledge graph induction system.
Extracting Common Inference Patterns from Semi-Structured Explanations (D19-60)

Copied to clipboard

Challenge: Multi-hop inference suffers from semantic drift, or the tendency for chains of reasoning to "drift"' to unrelated topics.
Approach: They propose to extract large high-confidence multi-hop inference patterns from a corpus of explanations by abstracting large-scale structure from logical sentences.
Outcome: The proposed method extracts large high-confidence multi-hop inference patterns from a “matter” subset of elementary science exam questions.
The DBpedia Databus Tutorial: Increase the Visibility and Usability of Your Data (2024.lrec-tutorials)

Copied to clipboard

Challenge: Linked Open Data tutorial introduces DBpedia Databus, a FAIR data publishing platform . aimed at addressing data production and consumption challenges faced by knowledge graph stakeholders .
Approach: This tutorial introduces DBpedia Databus, a FAIR data publishing platform . it addresses challenges faced by data producers and consumers .
Outcome: This tutorial addresses challenges faced by data producers and consumers using DBpedia Databus.
Knowledge Graphs meet Moral Values (2020.starsem-1)

Copied to clipboard

Challenge: Moral Foundations Theory (MFT) is one of the most adopted theories of morality due to its accompanying lexicon, the Moral Foundation Dictionary (MFD).
Approach: They propose to use the Moral Foundation Dictionary to analyze moral values in three widely used KGs and propose several Personalized PageRank variations to score concepts and entities in the KG with respect to their relevance to the different moral values.
Outcome: The proposed methods help to operationalize morality in both NLP and KG communities.
ELEVANT: A Fully Automatic Fine-Grained Entity Linking Evaluation and Analysis Tool (2022.emnlp-demos)

Copied to clipboard

Challenge: a tool for fine-grained evaluation of entity linkers on benchmarks is presented . a typical evaluation shows that overall precision and recall are poor . particular benchmarks often require very particular skills from an entity linker .
Approach: They propose a tool for fine-grained evaluation of entity linkers on benchmarks . they use a graph-based tool to analyze performance of a set of entity links .
Outcome: The proposed tool provides an automatic breakdown of the performance by error categories and by entity type.
Unified Examination of Entity Linking in Absence of Candidate Sets (2024.naacl-short)

Copied to clipboard

Challenge: Entity linking systems depend on candidate sets for their performance, but a comprehensive comparative analysis of these systems is lacking.
Approach: They propose a black-box benchmark and a method to evaluate all state-of-the-art entity linking methods.
Outcome: The proposed approach reduces the inference time and memory footprint of some models.
Identification of Alias Links among Participants in Narratives (P18-2)

Copied to clipboard

Challenge: Identifying distinct and independent participants in a narrative is crucial for many NLP applications.
Approach: They propose an approach based on linguistic knowledge for identification of aliases mentioned using proper nouns, pronouns or noun phrases with common noun headword.
Outcome: The proposed approach performs better than the state-of-the-art approach on four diverse history narratives of varying complexity.
A Web-scale system for scientific knowledge exploration (P18-4)

Copied to clipboard

Challenge: a system that organizes scientific knowledge into a hierarchical concept structure is needed to enable efficient exploration of Web-scale knowledge.
Approach: They propose a system that organizes scientific knowledge into a hierarchical concept structure . system allows researchers to identify hundreds of thousands of scientific concepts . it also allows researchers tagging scientific publications into millions of concepts based on text and graph structure based model .
Outcome: The proposed system builds the most comprehensive cross-domain scientific concept ontology published to date, with more than 200 thousand concepts and over one million relationships.
Multilingual Autoregressive Entity Linking (2022.tacl-1)

Copied to clipboard

Challenge: mGENRE is a sequence-to-sequence system for multilingual entity linking . mGenRE is used to solve language-specific mentions to a multilingual Knowledge Base .
Approach: They propose a sequence-to-sequence system for multilingual entity linking . they match language-specific mentions against a multilingual Knowledge Base (KB) mGENRE is a sequential system that predicts the name of the target entity token-by-token .
Outcome: The proposed system improves on three popular MEL benchmarks and shows improvements in accuracy.
An Environment for Relational Annotation of Political Debates (P19-3)

Copied to clipboard

Challenge: Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities.
Approach: They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors.
Outcome: The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors.
Improving Knowledge Base Construction from Robust Infobox Extraction (N19-2)

Copied to clipboard

Challenge: Existing knowledge bases are incomplete, resulting in poor answers and incompleteness.
Approach: They propose a method to extract Wikipedia infobox tables to populate an existing KB.
Outcome: The proposed method improves accuracy and completeness of the final KB significantly compared to DBpedia's baseline method.
Reasoning Over Paths via Knowledge Base Completion (D19-53)

Copied to clipboard

Challenge: Existing methods to predict missing links in knowledge graphs are lacking.
Approach: They propose a method to automatically rank paths between a source and target entity pair using a knowledge base completion model.
Outcome: The proposed method can rank and rank paths in biomedical knowledge graphs with a KBC model.
HOSMEL: A Hot-Swappable Modularized Entity Linking Toolkit for Chinese (2022.acl-demo)

Copied to clipboard

Challenge: Existing studies have explored the use of entity linking (EL) in downstream tasks.
Approach: They propose a modularized entity linking toolkit for easy task adaptation.
Outcome: The proposed toolkit achieves significantly better accuracy and less time and spaceconsumption than existing methods.
An ensemble CNN method for biomedical entity normalization (D19-57)

Copied to clipboard

Challenge: Named entity recognition (NER) and entity normalization (entity linking) are two fundamental natural language processing tasks to achieve entity normalizing.
Approach: They propose a CNN method that normalizes microbiology-related entities to concepts in standard dictionaries.
Outcome: The proposed method performs well in the BioNLP-OST19 shared task Bacteria Biotope.
Simple Augmentations of Logical Rules for Neuro-Symbolic Knowledge Graph Completion (2023.acl-short)

Copied to clipboard

Challenge: Recent studies show that high-quality rule sets struggle with high coverage.
Approach: They propose three simple augmentations to existing rule sets to improve results . they propose transforming rules to their abductive forms and generating equivalent rules that use inverse forms of constituent relations .
Outcome: The proposed methods achieve up to 7.1 pt MRR and 8.5 pT Hits@1 gains over using rules without augmentations.
BENNERD: A Neural Named Entity Linking System for COVID-19 (2020.emnlp-demos)

Copied to clipboard

Challenge: a biomedical entity linking system is available for COVID-19 research.
Approach: They propose a biomedical entity linking system that detects named enti- ties in text and links them to the UMLS knowledge base.
Outcome: The proposed system detects named enti- ties in text and links them to the unified medical language system (UMS) knowledge base entries.
A Reproducibility Study on Quantifying Language Similarity: The Impact of Missing Values in the URIEL Knowledge Base (2024.naacl-srw)

Copied to clipboard

Challenge: URIEL aggregates linguistic information for 4,005 languages and computes distances based on this information.
Approach: They propose to use a typological knowledge base to quantify language similarity to investigate URIEL's ambiguity in calculating language distances and handling missing values.
Outcome: The URIEL knowledge base does not provide information about typological features for 31% of the languages it represents, undermining the reliability of the database, especially on low-resource languages.
Joint Learning of Named Entity Recognition and Entity Linking (P19-2)

Copied to clipboard

Challenge: Named entity recognition and entity linking are two fundamentally related tasks . most approaches focus on the mention detection part, assuming the correct mentions have been detected .
Approach: They perform joint learning of named entity recognition and entity linking to leverage their relatedness.
Outcome: The proposed model achieves competitive results with the state-of-the-art in both NER and EL tasks.
Goodwill Hunting: Analyzing and Repurposing Off-the-Shelf Named Entity Linking Systems (2021.naacl-industry)

Copied to clipboard

Challenge: Named entity linking (NEL) is a preprocessing step in commercial systems . a small organization or individual could use an off-the-shelf system to accomplish the same objectives .
Approach: They examine how to repurpose off-the-shelf NEL systems to correct sport-related errors.
Outcome: The proposed model can improve sports question-answering accuracy by 25% . the proposed model is based on the best available model .
Learning Dynamic Context Augmentation for Global Entity Linking (D19-1)

Copied to clipboard

Challenge: Existing collective entity linking methods are expensive and often lack local context information.
Approach: They propose a dynamic context-augmented inference model that can be used to make collective inference.
Outcome: The proposed model can cope with different local EL models with different learning settings, base models, decision orders and attention mechanisms.
LOA: Logical Optimal Actions for Text-based Interaction Games (2021.acl-demo)

Copied to clipboard

Challenge: et al., 2019) have proposed a neuro-symbolic approach for reinforcement learning in non-simultaneous environments.
Approach: They propose an action decision architecture with a neuro-symbolic framework for natural language interaction games.
Outcome: The proposed framework provides an open-source implementation in Python for the reinforcement learning environment to facilitate an experiment for studying neuro-symbolic agents.
Entity Linking in the Job Market Domain (2024.findings-eacl)

Copied to clipboard

Challenge: In Natural Language Processing, entity linking (EL) has centered around Wikipedia, but yet remains underexplored for the job market domain.
Approach: They propose to use a bi-encoder and an autoregressive model to link fine-grained span-level skill mentions to a specific taxonomy entry to quantify labor market demands.
Outcome: The proposed model outperforms GENRE in strict evaluation, but performs better in loose evaluation.
Generating Questions from Wikidata Triples (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods for question generation from knowledge bases rely on extensive pre- and post-processing of the input triple.
Approach: They revisit KBQG using pre training, a new (triple, question) dataset and taking question type into account and provide a more extended KBqg dataset.
Outcome: The proposed approach outperforms existing methods in a standard and in 'zero-shot' setting.
Semantically-informed Hierarchical Event Modeling (2023.starsem-1)

Copied to clipboard

Challenge: Existing approaches to event modeling combine sequential latent variables with semantic ontological knowledge to improve representational capabilities.
Approach: They propose a doubly hierarchical semi-supervised event modeling framework that provides structural hierarchy while accounting for ontological hierarchy.
Outcome: The proposed model outperforms state-of-the-art models by 8.5% across two datasets and four metrics.
When ACE met KBP: End-to-End Evaluation of Knowledge Base Population with Component-level Annotation (L18-1)

Copied to clipboard

Challenge: Automating constructing a Knowledge Base from unstructured text is a goal of natural language processing.
Approach: They propose a method to evaluate a Knowledge Base population from unstructured text . they propose bootstrap resampling to provide statistical significance to the results .
Outcome: The proposed method uses component-level annotations to evaluate Cold Start KBP . it also uses bootstrap resampling to provide statistical significance to the results reported .
Extraction of Diagnostic Reasoning Relations for Clinical Knowledge Graphs (2022.acl-srw)

Copied to clipboard

Challenge: Existing methods for analyzing knowledge graphs focus on concept relations and clinical processes.
Approach: They propose to extract clinical knowledge graphs from a wiki and consumer health resource texts by using a clinical reasoning ontology.
Outcome: The proposed methods evaluate the correctness of extracted triples in the zero-shot setting.
KGxBoard: Explainable and Interactive Leaderboard for Evaluation of Knowledge Graph Completion Models (2022.emnlp-demos)

Copied to clipboard

Challenge: Knowledge Graphs (KGs) store information in the form of (head, predicate, tail)-triples.
Approach: They propose a framework for performing fine-grained evaluation on meaningful subsets of data.
Outcome: The proposed framework tests models on meaningful subsets of the data, which would have been impossible to detect with standard averaged single-score metrics.
DESCGEN: A Distantly Supervised Datasetfor Generating Entity Descriptions (2021.acl-long)

Copied to clipboard

Challenge: Short textual descriptions of entities provide summaries of their key attributes but generating entity descriptions can be challenging since information is scattered across multiple sources with varied content and style.
Approach: They propose to generate an entity summary description from 37K entities from Wikipedia and Fandom, paired with nine evidence documents on average.
Outcome: The proposed task is entity-centric, more abstractive, and covers a wide range of domains.
MOLEMAN: Mention-Only Linking of Entities with a Mention Annotation Network (2021.acl-short)

Copied to clipboard

Challenge: Existing approaches to entity linking represent each entity with a single vector, but instead use a contextualized mention-encoder that learns to place similar mentions of the same entity closer in vector space than mentions from different entities.
Approach: They propose an instance-based nearest neighbor approach to entity linking that allows for a contextualized mention-encoder to learn to place similar mentions of the same entity closer in vector space than mentions from different entities.
Outcome: The proposed approach outperforms all other systems on two multilingual benchmarks and is simpler to train and interpretable.
Hands-On Interactive Neuro-Symbolic NLP with DRaiL (2022.emnlp-demos)

Copied to clipboard

Challenge: Existing methods to enhance and correct NLP models require feedback from users.
Approach: They propose to enhance DRaiL with an easy to use Python interface that allows users to define, modify and augment models, as well as debug and visualize the predictions.
Outcome: The proposed framework supports predicting sentence and entity level moral sentiment in political tweets.
BLINK with Elasticsearch for Efficient Entity Linking in Business Conversations (2022.naacl-industry)

Copied to clipboard

Challenge: Existing systems that align textual mentions of entities to knowledge bases are difficult to deploy in production environments.
Approach: They propose a neural entity linking system that connects entities in business phone conversations to their corresponding Wikipedia and Wikidata entries.
Outcome: The proposed system improves inference speed and memory consumption while maintaining high accuracy.
A Study of the Importance of External Knowledge in the Named Entity Recognition Task (P18-2)

Copied to clipboard

Challenge: Existing studies have shown that external knowledge is important for Named Entity Recognition .
Approach: They propose a modular framework that divides knowledge into four categories according to depth . they show the effects when incrementally adding deeper knowledge .
Outcome: The proposed framework outperforms agnostic frameworks with more external knowledge . the proposed frameworks outperformed agrarian frameworks on two standard datasets .
Framing Named Entity Linking Error Types (L18-1)

Copied to clipboard

Challenge: Named Entity Linking (NEL) and relation extraction forms the backbone of Knowledge Base Population tasks.
Approach: They propose a taxonomy to frame common errors and apply it to four well-known Named Entity Linking systems.
Outcome: The proposed taxonomy was applied to four well-known Named Entity Linking systems on three gold standards.
Disentangled Action Recognition with Knowledge Bases (2022.naacl-main)

Copied to clipboard

Challenge: a new method for compositional action recognition is proposed to address the problem of zero-shot learning.
Approach: They propose a method to generalize compositional action recognition models to new verbs and nouns . they use knowledge graphs to extract disentangled feature representations for verbs, noun and type constraint .
Outcome: The proposed approach improves generalization ability of the compositional action recognition model to novel verbs and nouns that are unseen during training time.
entity-linkings: A Unified Library for Entity Linking (2026.eacl-demo)

Copied to clipboard

Challenge: Entity linking (EL) is the task of mapping named entities in text to canonical entries in a knowledge base.
Approach: They propose a unified library for using and developing entity linking systems . a strong emphasis is placed on usability, making it highly extensible .
Outcome: a new library aims to disambiguate named entities in text by mapping them to canonical entries in a knowledge base.
HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for retrieving information from a semi-structured knowledge base are struggling with hybrid questions.
Approach: They propose a retrieval method that leverages both textual and relational information from a semi-structured knowledge base to answer user questions.
Outcome: The proposed method surpasses all baselines on the STaRK benchmark and achieves significant performance gains.
Neuro-Symbolic Agentic Reinforcement Learning for Long-Term Original Character Companionship and Interaction (2026.acl-short)

Copied to clipboard

Challenge: Existing LLM-based agents that are optimized by prompting or supervised fine-tuning exhibit a generalization gap in long-horizon, socially rich interactions.
Approach: They propose a framework that formalizes OC companion agents’ interactions as a POMDP and decomposes the agent into three sub-policies optimized via closed-loop RL from AI feedback with verifiable rewards in a graph-constrained action space.
Outcome: The proposed framework formalizes OC companion agents’ interactions as a POMDP and decomposes the agent into three sub-policies (Router, Memory, and Persona) with verifiable rewards in a graph-constrained action space.
Neural Tensor Networks with Diagonal Slice Matrices (N18-1)

Copied to clipboard

Challenge: A large number of parameters can cause overfitting and a long training time for neural tensor networks (NTNs).
Approach: They propose two new parameter reduction techniques to reduce the number of parameters in an NTN without diminishing its expressiveness.
Outcome: The proposed models learn better and faster than the original (R)NTNs.
KGPool: Dynamic Knowledge Graph Context Selection for Relation Extraction (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for relation extraction (RE) use only expanded facts from the knowledge graph .
Approach: They propose a method for relation extraction from a single sentence . they use a neural network to expand the context with additional facts from the KG .
Outcome: The proposed method is more accurate than state-of-the-art methods on standard datasets.
The Lifecycle of “Facts”: A Survey of Social Bias in Knowledge Graphs (2022.aacl-main)

Copied to clipboard

Challenge: Knowledge graphs are used in a variety of downstream tasks and in hybrid AI systems.
Approach: They propose to examine the lifecycle of knowledge graphs with respect to bias influences.
Outcome: The proposed models are based on the lifecycle of knowledge graphs and their embedded versions . they show that the KGs manifest biases and propagate harmful prejudices .
K-pop and fake facts: from texts to smart alerting for maritime security (2023.acl-industry)

Copied to clipboard

Challenge: Maritime security requires full-time monitoring of the situation based on technical data but also from OSINT-like inputs.
Approach: They propose a system that extracts data from sensors and texts to feed a Knowledge Base . the system can be used to detect malicious actors using AIS and pseudo-newspapers .
Outcome: The proposed system ingests data from sensors and texts and feeds a Knowledge Base . it performs coherence checks between extracted facts and the extracted data .
From Superficial to Deep: Integrating External Knowledge for Follow-up Question Generation Using Knowledge Graph and LLM (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for generating follow-up questions are limited to shallow contextual questions that are uninspiring and have a large gap to the human level.
Approach: They propose a three-stage external knowledge-enhanced follow-up question generation method which generates questions by identifying contextual topics, building a knowledge graph online, and finally combining these with a large language model to generate the final question.
Outcome: The proposed method generates questions by identifying contextual topics, building a knowledge graph (KG) online, and finally combining these with a large language model to generate the final question.
Systematic Study of Long Tail Phenomena in Entity Linking (C18-1)

Copied to clipboard

Challenge: Existing systems for entity linking are based on frequent 'head' cases, while performance drops when moving towards rare 'long tail' entities.
Approach: They propose to use a long tail to investigate the properties of entity linking datasets.
Outcome: The proposed systems overfit to popular/frequent and non-ambiguous cases and find the most difficult cases among the infrequent candidates of ambiguous forms.
Neural Collective Entity Linking (C18-1)

Copied to clipboard

Challenge: Entity linking aims to link entity mentions in texts to knowledge bases, but existing methods rely on local contexts to resolve entities independently.
Approach: They propose a neural model for collective entity linking that integrates local contextual features and global coherence information to improve the computation efficiency.
Outcome: The proposed model improves its performance on five publicly available datasets and can be used to train on Wikipedia hyperlinks to avoid overfitting and domain bias.
CEO: Corpus-based Open-Domain Event Ontology Induction (2024.findings-eacl)

Copied to clipboard

Challenge: Existing event-centric NLP models restrict their generalization capabilities by limiting the pre-defined ontology.
Approach: They propose a Corpus-based Event Ontology induction model to relax the restriction imposed by pre-defined ontologies.
Outcome: The proposed model can induce a hierarchical event ontology with meaningful names on eleven open-domain corpora, making it more trustworthy and easier to be further curated.
Linking Entities to Unseen Knowledge Bases with Arbitrary Schemas (2021.naacl-main)

Copied to clipboard

Challenge: Existing work on entity linking relies on a knowledge base that is not known at training time.
Approach: They propose a method to flexibly convert entities with several attribute-value pairs from arbitrary KBs into flat strings and use it to generalize the model.
Outcome: The proposed model is 12% more accurate than baseline models on English datasets.
Fine-Grained Evaluation for Entity Linking (D19-1)

Copied to clipboard

Challenge: Entity Linking (EL) is an Information Extraction task that identifies entity mentions in a text corpus and associates them with an unambiguous identifier in KBs such as Wikipedia, BabelNet, DBpedia, Wikidata and YAGO.
Approach: They propose a fine-grained categorization of different types of entity mentions and links and propose 'fuzzy recall' metric to address the lack of consensus and compare a selection of online EL systems.
Outcome: The proposed task offers a bridge between unstructured text and structured KBs, where EL has applications for semantic search, document classification, relation extraction, and more.
Parallel Instance Query Network for Named Entity Recognition (2022.acl-long)

Copied to clipboard

Challenge: Named entity recognition is a fundamental task in natural language processing.
Approach: They propose a method that sets up global and learnable instance queries to extract entities from a sentence in a parallel manner.
Outcome: The proposed method outperforms existing state-of-the-art models on nested and flat datasets.
KnowledgeNet: A Benchmark Dataset for Knowledge Base Population (D19-1)

Copied to clipboard

Challenge: KnowledgeNet provides text exhaustively annotated with facts . high-quality KBs still rely almost exclusively on human-curated structured or semi-structured data.
Approach: They propose five baseline approaches to populating a knowledge base with facts . the best approach achieves an F1 score of 0.50, significantly outperforming a traditional approach by 79% .
Outcome: The best approach achieves an F1 score of 0.50, outperforming a traditional approach by 79%, indicating the dataset is challenging.
Knowledge-Enhanced Named Entity Disambiguation for Short Text (2020.aacl-main)

Copied to clipboard

Challenge: Existing methods for named entity disambiguation are weak for short text . performance of existing methods drops dramatically for short texts .
Approach: They propose a knowledge-enhanced method for named entity disambiguation . they use factual knowledge graph and conceptual knowledge graph to provide additional knowledge .
Outcome: The proposed method achieves significant improvement on a large manually annotated short-text dataset and the state-of-the-art on three standard datasets.
Temporal Knowledge Graph Reasoning Based on N-tuple Modeling (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing Temporal Knowledge Graphs (TKGs) only contain their core entities and form them as quadruples.
Approach: They propose to describe a temporal fact more accurately as an n-tuple . they propose to use a neural network to learn evolutional representations of entities .
Outcome: The proposed model oversimplifies and causes information loss on two datasets.
EventWiki: A Knowledge Base of Major Events (L18-1)

Copied to clipboard

Challenge: Existing knowledge bases focus on static entities such as people, locations and organizations.
Approach: They propose a new knowledge base resource called EventWiki which concentrates on major events . they show that EventWiki is a very useful resource for information extraction regarding events in NLP .
Outcome: The proposed resource is the first knowledge base resource of major events.
UAQFact: Evaluating Factual Knowledge Utilization of LLMs on Unanswerable Questions (2025.findings-acl)

Copied to clipboard

Challenge: Existing datasets to assess LLMs' performance on unanswerable questions lack factual knowledge support.
Approach: They propose a bilingual unanswerable question dataset with auxiliary factual knowledge created from a Knowledge Graph and two new tasks to measure LLMs' ability to utilize internal and external factual information.
Outcome: The proposed datasets show that LLMs do not consistently perform well even when they have factual knowledge stored.
CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation (2026.acl-long)

Copied to clipboard

Challenge: a hallmark of human innovation is recombination.
Approach: They propose a task to extract recombination instances from scientific literature . they analyze patterns of recombined concepts and apply it to a broad corpus of AI papers .
Outcome: The proposed model can predict cross-disciplinary research directions . it can predict recombinations across areas and link methods and concepts .
Advancing Event Causality Identification via Heuristic Semantic Dependency Inquiry Network (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for ECI rely on causal features and external knowledge, but these methods fail in two dimensions: causal features between events in texts often lack explicit clues and external information may introduce bias.
Approach: They propose a simple and effective Semantic Dependency Inquiry Network for ECI that captures semantic dependencies within the context using a unified encoder and generates a fill-in token based on comprehensive context understanding.
Outcome: Extensive experiments show that SemDI surpasses state-of-the-art methods on three widely used benchmarks.
Topic Ontologies for Arguments (2023.findings-eacl)

Copied to clipboard

Challenge: Many computational argumentation tasks, such as stance classification, are topic-dependent.
Approach: They map the argumentation landscape using the World Economic Forum, Wikipedia and Debatepedia as sources for argument topics.
Outcome: The argument ontology is the first comprehensive assessment of argument topics in argument corpora.
Guess Me if You Can: Acronym Disambiguation for Enterprises (P18-1)

Copied to clipboard

Challenge: Acronyms are abbreviations formed from the initial components of words or phrases . acronyms can be difficult to understand for people who are not familiar with the subject matter .
Approach: They propose a framework to automatically resolve the true meanings of acronyms in a given context . they use the enterprise corpus as input and a high-quality acronym disambiguation system as output .
Outcome: The proposed framework can be deployed to any enterprise to support acronym disambiguation.
Open Domain Question Answering based on Text Enhanced Knowledge Graph with Hyperedge Infusion (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve knowledge base are incomplete and difficult to understand.
Approach: They propose a novel QA method by leveraging text information to enhance the incomplete KB.
Outcome: Extensive experiments on the WebQuestionsSP benchmark prove the effectiveness of the proposed model.
KnowDis: Knowledge Enhanced Data Augmentation for Event Causality Detection via Distant Supervision (2020.coling-main)

Copied to clipboard

Challenge: Existing methods of event causality detection use hand-labeled training data.
Approach: They propose a framework for event causality detection that augments training data via distant supervision.
Outcome: The proposed framework outperforms existing methods on two benchmark datasets . it outperformed previous methods by a large margin assisted with automatically labeled training data.
Making a Semantic Event-type Ontology Multilingual (2022.lrec-1)

Copied to clipboard

Challenge: a new version of SynSemClass is being developed for use in natural language processing . the ontology is a bilingual resource with no links to a valency lexicon .
Approach: They propose to add German entries to the SynSemClass Event-type Ontology . they propose to use the ontology as a human-readable and human-understandable database .
Outcome: The proposed extension of SynSemClass Event-type Ontology is presented in a paper in czech republic . the ontology provides curated data for NLP experiments with cross-lingual synonyms .
Discovering Implicit Knowledge with Unary Relations (P18-1)

Copied to clipboard

Challenge: State-of-the-art relation extraction methods only recognize relationships between mentions of entity arguments stated explicitly in the text.
Approach: They propose a method to identify relations between two entities using unary relations and a common deep learning based representation.
Outcome: The proposed method outperforms state-of-the-art relation extraction technology on a web scale knowledge base population benchmark.
Improving Entity Linking by Modeling Latent Relations between Mentions (P18-1)

Copied to clipboard

Challenge: Entity linking systems often exploit relations between textual mentions to decide if the linking decisions are compatible.
Approach: They treat relations as latent variables while optimizing the neural entity-linking model without supervision.
Outcome: The proposed model outperforms its relation-agnostic version and significantly outperformed its relational version.
A Multi-Attention based Neural Network with External Knowledge for Story Ending Predicting Task (C18-1)

Copied to clipboard

Challenge: Existing studies on the topic of common sense story understanding focus on generating guesses for a missing event or concentrating on unsupervised learning.
Approach: They propose to extend attention-based neural network with external knowledge resources to understand temporal stories and predict their endings.
Outcome: The proposed model outperforms state-of-the-art models and external knowledge resources.
Counterfactual Supporting Facts Extraction for Explainable Medical Record Based Diagnosis with Graph Network (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods provide explanations based on a precise medical knowledge base, which is disease-specific and difficult to obtain for experts in reality.
Approach: They propose a method to extract supporting facts from irregular EMR without external knowledge bases by constructing a hierarchical graph network and using it to obtain causal relationship between multi-granularity features and diagnosis results.
Outcome: The proposed method diagnoses four types of EMR correctly and provides accurate supporting facts for the results.
End-to-End Construction of NLP Knowledge Graph (2021.findings-acl)

Copied to clipboard

Challenge: a new schema for NLP knowledge about tasks, datasets and metrics is proposed.
Approach: They propose a new schema that represents knowledge about tasks, datasets and metrics in the NLP domain.
Outcome: The proposed framework can be automatically built into scientific leaderboards . the proposed system achieves reasonable results for all relation types on this small-scale graph .
Find the Funding: Entity Linking with Incomplete Funding Knowledge Bases (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches to identifying and linking funding entities are suboptimal for the funding domain.
Approach: They propose an entity linking model that can perform NIL prediction and overcome data scarcity issues in a time and data-efficient manner.
Outcome: The proposed model outperforms existing baselines and overcomes data scarcity issues in a time and data-efficient manner.
Generation and Extraction Combined Dialogue State Tracking with Hierarchical Ontology Integration (2021.emnlp-main)

Copied to clipboard

Challenge: Current models are not satisfactory for solving out-of-vocabulary problems . current models assume that the task ontology is well defined in advance .
Approach: They propose to enhance the interrelation between slots with masked hierarchical attention.
Outcome: The proposed model yields a significant performance gain over current state-of-the-art model and is more robust to out-ofvocabulary problem compared with other methods.
ExCAR: Event Graph Knowledge Enhanced Explainable Causal Reasoning (2021.acl-long)

Copied to clipboard

Challenge: Existing work infers the causation between events based on knowledge from annotated causal event pairs, but additional evidence information is unexploited.
Approach: They propose an Event graph knowledge enhanced explainable CAusal Reasoning framework that acquires additional evidence information from a large-scale causal event graph as logical rules for causal reasoning.
Outcome: The proposed framework outperforms state-of-the-art methods in human evaluation and in animal models.
JUREX-4E: Juridical Expert-Annotated Four-Element Knowledge Base for Legal Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have introduced legal theories into LLM workflows to improve their understanding of legal texts and reasoning accuracy.
Approach: They evaluate an expert-annotated four-element knowledge base covering 155 criminal charges.
Outcome: The proposed model can be used to analyze criminal charges and retrieve them in legal cases.
Context-aware Entity Typing in Knowledge Graphs (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for knowledge graph entity typing are embedding-based and graph convolutional networks (GCNs) . Existing approaches for knowledge Graph Entity Typing (KGET) are incomplete and require multiple inference mechanisms.
Approach: They propose a method that uses entities’ contextual information to infer missing types in knowledge graphs by using two inference mechanisms: N2T and Agg2T.
Outcome: The proposed method can infer entities' missing types by completing two real-world KGs.
Clustering-based Inference for Biomedical Entity Linking (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches to linking entities ignore relationships between entities in biomedical knowledge bases.
Approach: They propose a model which can link mentions of unseen entities using learned representations of entities.
Outcome: The proposed model improves on the largest publicly available biomedical dataset by 3.0 points of accuracy and 2.3 points of reliability.
Improving Entity Disambiguation by Reasoning over a Knowledge Base (2022.naacl-main)

Copied to clipboard

Challenge: Recent work in entity disambiguation relies on a limited subset of KB facts to link entities . less common entities are prone to missing or inconsistent KB information, which is problematic for models which rely on 'one source'
Approach: They propose an ED model which links entities by reasoning over a symbolic knowledge base in a fully differentiable fashion.
Outcome: The proposed model outperforms state-of-the-art models on six well-established datasets by 1.3 F1 on average.
Reasoning with Trees: Faithful Question Answering over Knowledge Graph (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have shown remarkable progress in reasoning capabilities, yet they still face challenges in complex, multi-step reasoning tasks.
Approach: They propose a framework that synergistically integrates LLMs with knowledge graphs (KGs) to enhance reasoning performance and interpretability.
Outcome: The proposed framework outperforms existing state-of-the-art methods on two benchmark KGQA datasets and improves on the MCTS process.
Biomedical Concept Normalization over Nested Entities with Partial UMLS Terminology in Russian (2024.lrec-main)

Copied to clipboard

Challenge: Existing annotations in Russian do not include all entities, but only a small fraction of them are labeled in English.
Approach: They present a manually annotated PubMed abstract dataset for concept normalization in Russian.
Outcome: The proposed model improves on nested named entities in a zero-shot setting on bilingual terminology.
Dynamic Topic Tracker for KB-to-Text Generation (2020.coling-main)

Copied to clipboard

Challenge: Existing KB-to-text generation models suffer from an off-topic problem . existing models generate unrelated clauses regardless of input data .
Approach: They propose a dynamic topic tracker that learns a global hidden representation for topics and recognizes the corresponding topic during each generation step.
Outcome: The proposed model improves the performance of sentence generation and mitigates off-topic problem.
External Knowledge-Driven Argument Mining: Leveraging Attention-Enhanced Multi-Network Models (2024.emnlp-main)

Copied to clipboard

Challenge: Argument mining involves the identification of argument relations (AR) between Argumentative Discourse Units (ADUs).
Approach: They propose to leverage external resources to identify semantic paths linking ADUs . they propose to use WordNet, ConceptNet, and Wikipedia to identify these paths .
Outcome: The proposed architecture achieves F-scores of 0.85, 0.84, 0.70, and 0.87 on four datasets.
Named Entity Recognition for Entity Linking: What Works and What’s Next (2021.findings-emnlp)

Copied to clipboard

Challenge: Entity Linking (EL) systems have achieved impressive results on standard benchmarks thanks to the contextualized representations provided by recent pretrained language models.
Approach: They propose to exploit Named Entity Recognition (NER) to narrow the gap between EL systems trained on high and low amounts of labeled data.
Outcome: The proposed model can be exploited to narrow the gap between EL systems trained on high and low amounts of labeled data.
Knowledge Enhanced Reflection Generation for Counseling Dialogues (2022.acl-long)

Copied to clipboard

Challenge: Using retrieval and generative methods, we generate responses using commonsense and domain knowledge.
Approach: They propose a pipeline that collects domain knowledge through web mining and a model that incorporates knowledge generated by COMET using soft positional encoding and masked self-attention.
Outcome: The proposed pipeline collects domain knowledge through web mining and incorporates knowledge generated by COMET using soft positional encoding and masked self-attention.
How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on hierarchical organization of categories focused on basic-1 . but, words at the subordinate level are crucial for effective communication in specialized domains.
Approach: They analyze a psycholinguistic dataset of human-generated exemplars for 187 concrete words . they then evaluate whether textual and vision LLMs produce meaningful exemplar .
Outcome: The results show that human-generated exemplars perform poorly in three key tasks . the results highlight the potential of using AI-generated categories in psycholinguistic research .
HTCCN: Temporal Causal Convolutional Networks with Hawkes Process for Extrapolation Reasoning in Temporal Knowledge Graphs (2024.naacl-long)

Copied to clipboard

Challenge: Temporal knowledge graphs (TKGs) are powerful tools for storing and modeling dynamic facts.
Approach: They propose a Hawkes process-based temporal causal convolutional network for temporal reasoning under extrapolation settings.
Outcome: The proposed network is based on Hawkes process-based temporal causal convolutional network and captures the temporal evolution of facts.
Entity Linking within a Social Media Platform: A Case Study on Yelp (D18-1)

Copied to clipboard

Challenge: Existing studies on entity linking focus on linking entities to knowledge bases, but on social media platforms, such as Yelp, it can be more practical.
Approach: They propose to link entities within a social media platform with a new entity linking problem.
Outcome: The proposed model can link business mentions to corresponding businesses on a social media platform.
ZiNet: Linking Chinese Characters Spanning Three Thousand Years (2022.findings-acl)

Copied to clipboard

Challenge: tens of thousands of ancient characters must be deciphered by experts to interpret unearthed documents.
Approach: They propose a diachronic Chinese knowledge base to help researchers discover glyph similar characters by measuring glyph similarities between ancient Chinese characters.
Outcome: The proposed method shows strong correlations between the scores obtained from the method and from human experts.
Laying the Groundwork for Knowledge Base Population: Nine Years of Linguistic Resources for TAC KBP (L18-1)

Copied to clipboard

Challenge: Knowledge Base Population (KBP) evaluations target information extraction technologies for knowledge bases comprised of entities, relations, and events.
Approach: They describe the linguistic resources provided by Linguistic Data Consortium for TAC KBP since 2009 . they highlight changes made to support evolving evaluation requirements .
Outcome: The evaluations have targeted information extraction technologies for the population of knowledge bases comprised of entities, relations, and events.
Generating Questions for Knowledge Bases via Incorporating Diversified Contexts and Answer-Aware Loss (D19-1)

Copied to clipboard

Challenge: Conventional methods for question generation neglect two crucial research issues: 1) the given predicate needs to be expressed; 2) the answer to the generated question needs to have a definitive answer.
Approach: They propose a neural encoder-decoder model with multi-level copy mechanisms to generate questions . they also introduce answer-aware loss to make generated questions correspond to more definitive answers.
Outcome: The proposed model achieves state-of-the-art performance while corresponding to more definitive answers.
CLEEK: A Chinese Long-text Corpus for Entity Linking (2020.lrec-1)

Copied to clipboard

Challenge: Entity linking is a fundamental task in natural language processing, says nigel kilgstrom . existing corpora for entity linking in china are lacking and deficient, he says . kilsmstrom: a new method for entity disambiguation can be developed for Chinese .
Approach: They build a Chinese corpus of multi-domain long text for entity linking . they evaluate the difficulty of documents with respect to entity linking using a measure .
Outcome: The proposed corpus is based on 100 documents from diverse domains and is publicly accessible.
COMETA: A Corpus for Medical Entity Linking in the Social Media (2020.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for Entity Linking (EL) fail to address the complex nature of health terminology in layman’s language.
Approach: They propose to use a corpus of 20k English biomedical entity mentions from Reddit expert-annotated with links to a widely-used medical knowledge graph to investigate the ability of these systems to perform complex inference on entities and concepts.
Outcome: The proposed corpus satisfies a combination of desirable properties, from scale and coverage to diversity and quality, that to the best of our knowledge has not been met by existing resources in the field.
Movie101: A New Movie Understanding Benchmark (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to narrate movies with no actors are difficult to implement in real situations . a new metric is proposed to provide the best correlation with human evaluation .
Approach: They propose a large-scale Chinese movie benchmark to help visually impaired enjoy movies . they propose metric called Movie Narration Score (MNScore) which achieves best correlation with human evaluation.
Outcome: The proposed method outperforms baselines and the existing methods.
Conformal Event Prediction with Temporal Knowledge Graph (2026.findings-acl)

Copied to clipboard

Challenge: Current event prediction methods lack rigorous uncertainty quantification, which limits their reliability for decision-making.
Approach: They propose a conformal prediction framework that applies conformal predictions to event prediction to address this challenge.
Outcome: The proposed framework guarantees coverage while improving efficiency on three public datasets.
DIVINE: A Generative Adversarial Imitation Learning Framework for Knowledge Graph Reasoning (D19-1)

Copied to clipboard

Challenge: Existing knowledge graph reasoning methods require numerous trials for path-finding and require meticulous reward engineering to fit specific datasets.
Approach: They propose a plug-and-play framework that uses generative adversarial imitation learning to enhance existing RL-based methods.
Outcome: The proposed framework improves existing RL-based methods while eliminating reward engineering.
NeuSTIP: A Neuro-Symbolic Model for Link and Time Prediction in Temporal Knowledge Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Temporal Knowledge Graphs (KGs) are factual information repositories where a fact is associated with a time interval.
Approach: They propose a temporal NS model for knowledge graph completion that performs link prediction and time interval prediction in a TKG.
Outcome: The proposed model shows competitive performance on link prediction and time prediction.
Humans Keep It One Hundred: an Overview of AI Journey (2020.lrec-1)

Copied to clipboard

Challenge: Artificial General Intelligence (AGI) is showing growing performance in numerous applications - beating human performance in Chess and Go, using knowledge bases and text sources to answer questions and even pass human examination.
Approach: They propose to use knowledge bases and text sources to answer questions to improve AI performance on knowledge bases, reasoning and text generation.
Outcome: The proposed AI Journey system passed the final native language exam in Russian with a high score of 69%, with 68% being an average human result.
Pattern-revising Enhanced Simple Question Answering over Knowledge Bases (C18-1)

Copied to clipboard

Challenge: Simple question answering over knowledge bases is one of the most important natural language processing tasks.
Approach: They propose to conduct pattern extraction and entity linking first and put forward pattern revising procedure to mitigate the error propagation problem.
Outcome: The proposed method outperforms the current state-of-the-art in this task by an absolute large margin.
Orchestrating NLP Services for the Legal Domain (2020.lrec-1)

Copied to clipboard

Challenge: a legal technology system under development in the EU is based on semantic services and a multilingual legal knowledge Graph.
Approach: They propose a workflow manager that enables flexible orchestration of workflows . they describe different use cases and propose prototypical solutions .
Outcome: The proposed system is based on a set of natural language processing and document curation services and a multilingual legal knowledge graph that contains semantic information and meaningful references to legal documents.
Evaluation Dataset and Methodology for Extracting Application-Specific Taxonomies from the Wikipedia Knowledge Graph (2020.lrec-1)

Copied to clipboard

Challenge: Recent efforts to extract hierarchical relations from unstructured text have been challenging.
Approach: They propose an iterative method to extract an application-specific gold standard dataset from a Wikipedia knowledge graph and an evaluation framework to assess the quality of noisy automatically extracted taxonomies.
Outcome: The proposed method reduces manual work and provides a first gold standard dataset and evaluation framework.
Hierarchical User Intent Inference with Knowledge Graph Grounding (2026.findings-eacl)

Copied to clipboard

Challenge: Existing large language models lack structured grounding and do not capture nuanced intent expression.
Approach: They propose a Hierarchical Intent Inference framework that first predicts fine-grained aspect ratings and then generates natural language intent statements guided by contextual subgraphs retrieved from a domain-specific knowledge graph.
Outcome: The proposed framework outperforms strong LLM and encoder-based baselines on a hotel review dataset.
Representing Multiword Term Variation in a Terminological Knowledge Base: a Corpus-Based Study (2020.lrec-1)

Copied to clipboard

Challenge: Multiword terms are the most frequent type of lexical units in scientific and technical communication. rendering them in another language is not easy due to their cognitive complexity, proliferation of different forms, and their unsystematic representation in terminographic resources.
Approach: They evaluated Spanish translation variants of multiword terms in three parallel corpora, two comparable corporales and two terminological resources.
Outcome: The results show that multiword terms exhibit a significant degree of term variation . the proposed model is based on a set of criteria for determining which variants should be selected .
KORE 50ˆDYWC: An Evaluation Data Set for Entity Linking Based on DBpedia, YAGO, Wikidata, and Crunchbase (2020.lrec-1)

Copied to clipboard

Challenge: A major domain of research in natural language processing is named entity recognition and disambiguation (NERD).
Approach: They extend a widely-used data set to include NERD tasks for DBpedia and YAGO, Wikidata and Crunchbase.
Outcome: The extended data set allows for a broader spectrum of evaluation.
Improving Candidate Retrieval with Entity Profile Generation for Wikidata Entity Linking (2022.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on Wikipedia-derived KBs, but there is little work on EL over Wikidata . EL systems have found applications in many tasks such as question answering .
Approach: They propose a novel approach to linking entity mentions to referent entities in a knowledge base . they use a sequence-to-sequence model to generate the profile of the target entity .
Outcome: The proposed approach achieves state-of-the-art results on three Wikidata-based datasets and strong performance on TACKBP-2010.
Fact Discovery from Knowledge Base via Facet Decomposition (N19-1)

Copied to clipboard

Challenge: Recent years have witnessed the emergence and growth of many large-scale knowledge bases (KBs) however, there are some issues unsettled towards enriching the KBs.
Approach: They propose a framework that decomposes the discovery problem into several facet components and an auto-encoder component to estimate some facets of the fact.
Outcome: The proposed framework achieves promising results on a benchmark dataset.
KAPA: A Deliberative Agent Framework with Tree-Structured Knowledge Base for Multi-Domain User Intent Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on the use of LLMs for estimating user intents are either too far from real human thought processes or require labeled samples.
Approach: They propose a deliberative agent framework that leverages human thought process to build high-level domain knowledge and a tree-structured knowledge base to store refined experience and data.
Outcome: The proposed framework is able to build high-level domain knowledge and efficiently store it across multiple steps.
Thinking Beyond the Local: Multi-View Instructed Adaptive Reasoning in KG-Enhanced LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for large language models adopt query-driven iterative reasoning from a local perspective, limiting efficiency and accuracy for complex multi-hop tasks.
Approach: They propose a multi-view instructed adaptive reasoning of LLM on Knowledge Graphs that allows LLMs to plan, evaluate, and adapt reasoning paths from a global perspective.
Outcome: The proposed model overcomes the limitations of local exploration by enabling LLMs to plan, evaluate, and adapt reasoning paths from a global perspective.
WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity Types (2022.acl-long)

Copied to clipboard

Challenge: Multimodal Entity Linking (MEL) is an essential task for many multimodal applications.
Approach: They propose to use a human-annotated Wikipedia-based multimodal entity linking dataset to improve the quality of existing MEL models.
Outcome: The proposed model uses the visual information of images more effectively than existing models.
Joint Multilingual Knowledge Graph Completion and Alignment (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing work on multilingual KG completion has focused on entity and relation alignments, but understanding of how it can aid multilingual alignments is limited.
Approach: They propose to combine two components that jointly accomplish KG completion and alignment.
Outcome: The proposed model outperforms existing competitive baselines on a public multilingual benchmark and achieves state-of-the-art results.
Controllable Contrastive Generation for Multilingual Biomedical Entity Linking (2023.emnlp-main)

Copied to clipboard

Challenge: Multilingual biomedical entity linking (MBEL) aims to map language-specific mentions in biomedically text to standardized concepts in a multilingual knowledge base (KB).
Approach: They propose a prompt-based controllable contrastive generation framework for MBEL which summarizes multidimensional information of the UMLS concept mentioned in biomedical text into a natural sentence following a predefined template.
Outcome: The proposed framework matches against UMLS concepts in as many languages and types as possible, thus facilitating cross-information disambiguation.
MCMH: Learning Multi-Chain Multi-Hop Rules for Knowledge Graph Reasoning (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing work on knowledge graphs infers a missing relationship between entities with a multi-hop rule . Empirical results show that our multi-chain multi-homing (MCMH) rules yield superior results compared to the standard single-chain approaches.
Approach: They propose to use a generalized form of multi-hop rules to learn generalized rules efficiently . they propose to select a small set of relation chains as a rule and evaluate confidence .
Outcome: The proposed method outperforms the existing methods and the existing frameworks.
Extracting a Knowledge Base of Mechanisms from COVID-19 Papers (2021.naacl-main)

Copied to clipboard

Challenge: COVID-19 has spawned a diverse body of scientific literature that is challenging to navigate . researchers are using automated tools to help find useful knowledge .
Approach: They develop a schema to extract mechanism relations from scientific papers . their search engine, dataset and code are publicly available .
Outcome: The proposed schema outperforms PubMed search in clinical trials.
How Knowledge Graph and Attention Help? A Qualitative Analysis into Bag-level Relation Extraction (2021.acl-long)

Copied to clipboard

Challenge: Knowledge Graph (KG) and attention mechanism have been demonstrated effective in introducing and selecting useful information for weakly supervised methods.
Approach: They propose a paradigm to quantitatively evaluate the effect of attention and KG on bag-level relation extraction (RE) they propose to incorporate entity prior to KG-enhanced attention to improve RE performance .
Outcome: The proposed model achieves significant improvements on two real-world datasets compared with three state-of-the-art baselines.
The Integration of Semantic and Structural Knowledge in Knowledge Graph Entity Typing (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods to predict missing type annotations for knowledge graphs use only structural knowledge in the local neighborhood of entities.
Approach: They propose a model for KG Entity Typing that integrates semantic and structural knowledge to infer missing types.
Outcome: The proposed framework outperforms existing state-of-the-art methods in the Knowledge Graph Entity Typing task.
What do Entity-Centric Models Learn? Insights from Entity Linking in Multi-Party Dialogue (N19-1)

Copied to clipboard

Challenge: a recent study suggests that models that incorporate a bias towards learning entity representations are not effective at modeling entities.
Approach: They propose to use two entity-centric models for a referential task . they show they outperform the state of the art and do better on lower frequency entities .
Outcome: The proposed models outperform the state of the art on a referential task . they do better on lower frequency entities than a counterpart model not entity-centric .
Uncovering Probabilistic Implications in Typological Knowledge Bases (P19-1)

Copied to clipboard

Challenge: linguistic typology is concerned with mapping out the relationships between languages with structural and functional properties.
Approach: They propose a computational model which identifies known and new linguistic universals and uncovers them worthy of further linguistic investigation.
Outcome: The proposed model outperforms baselines and knowledge base baselines.
Multilingual Knowledge Editing with Language-Agnostic Factual Neurons (2025.coling-main)

Copied to clipboard

Challenge: Existing methods to update factual knowledge overlook connections of same knowledge between different languages, resulting in knowledge conflicts and limited edit performance.
Approach: They propose a method to edit multilingual knowledge simultaneously that avoids knowledge conflicts and improves edit performance.
Outcome: The proposed method avoids knowledge conflicts and improves edit performance on bi-ZsRE and MzsRE benchmarks.
The LODeXporter: Flexible Generation of Linked Open Data Triples from NLP Frameworks for Automatic Knowledge Base Construction (L18-1)

Copied to clipboard

Challenge: Linked Open Data (LOD) principles are used to export natural language processing (NLP) results to graph-based knowledge base.
Approach: They propose a method for exporting NLP results to a graph-based knowledge base using Linked Open Data principles.
Outcome: The proposed method is available as an open source component for the GATE framework and is available on GitHub.
ClinIDMap: Towards a Clinical IDs Mapping for Data Interoperability (2022.lrec-1)

Copied to clipboard

Challenge: a tool for mapping identifiers between clinical ontologies and lexical resources is available.
Approach: They propose a tool for mapping identifiers between clinical ontologies and lexical resources.
Outcome: The proposed mapping tool can be used to enrich existing annotated corpora in multiple languages.
Distant Learning for Entity Linking with Automatic Noise Detection (P19-1)

Copied to clipboard

Challenge: Accurate entity linkers have been produced for domains and languages where no or very limited amounts of labeled data are available.
Approach: They propose to use annotated text to learn to link entities without labeling . they frame the task as a multi-instance learning problem and rely on surface matching to create initial noisy labels.
Outcome: The proposed method outperforms the baseline surface matching model for a subset of entities.
Modeling Event-Pair Relations in External Knowledge Graphs for Script Reasoning (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods focus on graph triples with event overlap, but ignore more supportive triples . Script reasoning relies on understanding the relationship between two events .
Approach: They propose a model to learn the inferential relations between events from the whole eventuality KG . they propose 'script adapter' to extend the model to infer the associated relations between an event chain and a subsequent event candidate.
Outcome: The proposed model is compared with baselines using external KG or not on a script reasoning task.
A Multi-label Multi-hop Relation Detection Model based on Relation-aware Sequence Generation (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods treat multi-label learning problem as a single label . Existing approaches focus on measuring semantic similarity of questions and candidate relations .
Approach: They propose to solve multi-hop relation detection problem by generating sequences of hops and labels.
Outcome: The proposed method is effective in KBQA, despite the unknown number of labels and hops.
Corpus Query Lingua Franca part II: Ontology (2020.lrec-1)

Copied to clipboard

Challenge: outlines the projected second part of the Corpus Query Lingua Franca (CQLF) family of standards . the existence of a large number of different corpus query languages poses an epistemic challenge for the research community .
Approach: They propose to standardize the Corpus Query Lingua Franca (CQLF) family of standards . they present the assumptions and aims of the CQLF Metamodel and its basic structure .
Outcome: The proposed second part of the Corpus Query Lingua Franca (CQLF) family is in the process of standardization at the International Standards Organization (ISO) the first part of CQLF Ontology was adopted as an international standard at the beginning of 2018 .
A Fair and In-Depth Evaluation of Existing End-to-End Entity Linking Systems (2023.emnlp-main)

Copied to clipboard

Challenge: Existing evaluations of entity linking systems often lack detailed error analysis or a closer look at the results.
Approach: They evaluate existing entity linking systems and propose two new benchmarks . they characterize their strengths and weaknesses and report on reproducibility aspects .
Outcome: The evaluations of existing system have strong biases and artifacts . they characterize their strengths and weaknesses and report on reproducibility aspects .
Course Concept Expansion in MOOCs with External Knowledge and Interactive Game (P19-1)

Copied to clipboard

Challenge: Existing methods to expand course concepts in MOOCs suffer from semantic drifts and lack of knowledge guidance.
Approach: They propose to use a boundary search method to search for new concepts via external knowledge base and then use heterogeneous features to verify the results.
Outcome: The proposed method improves on the datasets from Coursera and XuetangX.
ChemNER: Fine-Grained Chemistry Named Entity Recognition with Ontology-Guided Distant Supervision (2021.emnlp-main)

Copied to clipboard

Challenge: Named entity recognition (NER) is a fundamental step in scientific literature analysis to build AI-driven systems for molecular discovery, synthetic strategy designing, and manufacturing.
Approach: They propose an ontology-guided method for fine-grained named entity recognition (NER) it leverages the chemistry type ontologies to generate distant labels with flexible KB-matching .
Outcome: The proposed method significantly outperforms the state-of-the-art methods with a .25 absolute F1 improvement.
Chinese Relation Extraction with Multi-Grained Information and External Linguistic Knowledge (P19-1)

Copied to clipboard

Challenge: Existing methods for Chinese relation extraction suffer from segmentation errors and ambiguity of polysemy.
Approach: They propose a multi-grained lattice framework for Chinese relation extraction . they incorporate word-level information into character sequence inputs to avoid segmentation errors .
Outcome: The proposed model outperforms existing models on three real-world datasets in distinct domains.
Inductive Relation Inference of Knowledge Graph Enhanced by Ontology Information (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to inference knowledge graphs lack ontology information, which is often too sparse.
Approach: They propose a knowledge graph inductive inference method that fuses ontology information to learn the semantic information of entities.
Outcome: The proposed method outperforms large language models like ChatGPT on two benchmark datasets and improves the MRR metrics by 15.4% and 44.1%, respectively.
Fine-grained Entity Typing without Knowledge Base (2021.emnlp-main)

Copied to clipboard

Challenge: Existing work on fine-grained entity typing (FET) relies on knowledge bases as distant supervision, but lack of or incompleteness of KB can hinder training.
Approach: They propose a two-step framework that trains FET models without accessing any knowledge base.
Outcome: The proposed framework achieves competitive performance with respect to the models trained on the original KB-supervised datasets.
Linking Surface Facts to Large-Scale Knowledge Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Open Information Extraction (OIE) methods extract facts in the form of triples . ambiguity of these triples hinders their downstream usage .
Approach: They propose a benchmark that measures fact linking performance on a granular triple slot level . they propose to use a system that can detect out-of-KG entities and predicates .
Outcome: The proposed benchmark can measure fact linking performance on a granular triple slot level while also measuring if a system can recognize that a surface form has no match in the existing KG.
Jointly Extracting Explicit and Implicit Relational Triples with Reasoning Pattern Enhanced Binary Pointer Network (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for relational triple extraction ignore implicit triples that lack explicit expressions, leading to incomplete knowledge graphs.
Approach: They propose a binary pointer network to extract explicit and implicit relational triples from sentences and to retain the information of extracted triples in an external memory.
Outcome: The proposed framework extracts overlapping triples relevant to each word sequentially and retains the information of extracted triples in an external memory.
Comparative evaluation of boundary-relaxed annotation for Entity Linking performance (2023.acl-long)

Copied to clipboard

Challenge: Entity Linking is a critical step for information extraction, allowing the retrieval and understanding of information from unstructured textual sources.
Approach: They propose to use noisy datasets to generate noisy versions of annotated entity mentions and then train three Entity Linking models on this data.
Outcome: The proposed model can be used to associate NE mentions to a single concept in an ontology, allowing for better indexing and relation extraction.
KG-Agent: An Efficient Autonomous Agent Framework for Complex Reasoning over Knowledge Graph (2025.acl-long)

Copied to clipboard

Challenge: Existing methods to design the interaction strategy between large language models and knowledge graphs (KGs) are not effective for large language model (LLM)s to solve complex tasks due to the large volume and structured format of KG data.
Approach: They propose an LLM-based agent framework that enables small LLMs to actively make decisions over knowledge graphs.
Outcome: The proposed framework outperforms existing methods on in-domain and out-domain datasets using 10K samples.
Entity Linking over Nested Named Entities for Russian (2022.lrec-1)

Copied to clipboard

Challenge: Entity linking is a popular NLP task, where a system needs to link a named entity to a concept in a knowledge base such as Wikidata.
Approach: They describe the main design principles behind entity linking annotation in the recently released Russian NEREL dataset for information extraction.
Outcome: The NEREL dataset is the largest Russian dataset annotated with entities and relations.
CONTOR: Benchmarking Strategies for Completing Ontologies with Plausible Missing Rules (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations focus on distinguishing held-out ontologies from randomly corrupted ones, which often makes the task unrealistically easy.
Approach: They propose to use the common description logic syntax for encoding ontology rules to test their effectiveness on manually annotated hard negatives.
Outcome: The proposed models are compared with existing models and have been evaluated on different ontologies.
Joint Biomedical Entity and Relation Extraction with Knowledge-Enhanced Collective Inference (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for information extraction from biomedical texts do not utilize external knowledge . despite the exponential growth of biomedically published articles, many existing methods fall behind .
Approach: They propose a framework that utilizes external knowledge for entity and relation extraction . KECI uses an initial span graph to construct a knowledge graph containing relevant background knowledge .
Outcome: The proposed framework achieves state-of-the-art results in two biomedical datasets . it achieves 4.59% and 4.91% improvement in F1 scores over the state- of-the art methods .
Reasoning with Ontology Graph: Toward Type-Constrained Knowledge Graph Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Existing knowledge graph question answering methods rely on LLM-induced type systems with inconsistent granularity or perform multi-hop reasoning without explicit target-type constraints.
Approach: They propose a type-constrained knowledge graph question answering framework that reasons over a relation-centric ontology graph.
Outcome: The proposed framework achieves state-of-the-art and produces ontology-grounded reasoning chains with substantial Hit@1 gains.
IntKB: A Verifiable Interactive Framework for Knowledge Base Completion (2020.coling-main)

Copied to clipboard

Challenge: Knowledge bases (KBs) present databases that store information about entities and relations among them.
Approach: They propose a question-based interactive framework for KB completion from text . their framework generates facts that are aligned with text snippets and is immediately verifiable by humans .
Outcome: The proposed framework achieves a hit@1 ratio of 29.7% for initial unseen relations, and gradually improves to 46.2%.
History repeats: Overcoming catastrophic forgetting for event-centric temporal knowledge graph completion (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for knowledge graph completion are incomplete and can lead to errors . retraining the model with the entire updated TKG can mitigate forgetting but is computationally burdensome.
Approach: They propose a temporal regularization framework that allows repurposing of parameters . they propose 'clustering-based experience replay' that reinforces the past knowledge .
Outcome: The proposed framework adapts to new events while reducing catastrophic forgetting.
Enhancing Event Causality Identification with LLM Knowledge and Concept-Level Event Relations (2025.coling-main)

Copied to clipboard

Challenge: Existing methods to identify causal relationships between events often overlook the dependencies between similar events.
Approach: They propose an ECI method enhanced by LLM Knowledge and Concept-Level Event Relations (LKCER) the method constructs a conceptual-level heterogeneous event graph by leveraging local contextual information of related event mentions.
Outcome: The proposed method outperforms previous state-of-the-art methods on both benchmarks, EventStoryLine and Causal-TimeBank.
Enhancing Biomedical Lay Summarisation with External Knowledge Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to lay summarisation are reliant on the source article, which is unlikely to include all the information necessary for a lay audience.
Approach: They augment existing biomedical lay summarisation dataset with article-specific knowledge graphs that contain detailed information on relevant biomedically related concepts.
Outcome: The proposed methods improve readability and explanation of technical concepts by integrating graph-based domain knowledge within lay summarisation models.
Controllable Open-ended Question Generation with A New Question Type Ontology (2021.acl-long)

Copied to clipboard

Challenge: Existing question types are limited to generating multiple-sense questions . we present a question type-aware question generation framework to generate open-ended questions based on multiple-phrase questions - a task that is less explored .
Approach: They propose a question type-aware question generation framework which predicts question focuses and produces the question.
Outcome: The proposed model improves question quality over competitive comparisons on large-scale datasets.
A Template Is All You Meme (2025.naacl-long)

Copied to clipboard

Challenge: Templatic memes are a form of communication capable of succinctly conveying complicated messages.
Approach: They propose a method to match memes to a knowledge base of 5,200 meme templates and 54,000 examples of template instances using a distance-based lookup.
Outcome: The proposed method improves general meme knowledge and sample efficiency, leading to more robust models.
Building a Multimodal Entity Linking Dataset From Tweets (2020.lrec-1)

Copied to clipboard

Challenge: Entity linking is a task that aims at associating an entity mention with a unique entity in a knowledge base.
Approach: They propose a method to quasi-automatically build annotated datasets to evaluate methods on the Entity Linking task.
Outcome: The proposed method builds annotated datasets of tweets with ambiguous mentions and a Twitter KB defining the entities.
Are Missing Links Predictable? An Inferential Benchmark for Knowledge Graph Completion (2021.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for Knowledge Graph Completion (KGC) are unsatisfactory .
Approach: They propose to use rule-guided train/test generation instead of conventional random split to ensure that each testing sample is predictable with supportive data in the training set.
Outcome: The proposed model improves on existing benchmarks in inferential ability, assumptions, and patterns.
Recurrent Event Network: Autoregressive Structure Inferenceover Temporal Knowledge Graphs (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for reasoning over temporal knowledge graphs focus on past timestamps and are not able to predict future interactions.
Approach: They propose a novel autoregressive architecture for predicting future interactions using a recurrent event encoder and a neighborhood aggregator.
Outcome: The proposed method achieves state-of-the-art on five public datasets.
Natural Evolution-based Dual-Level Aggregation for Temporal Knowledge Graph Reasoning (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models ignore asynchronous characteristics of event evolution, resulting in suboptimal performance.
Approach: They propose a Natural Evolution-based Dual-level Aggregation framework for TKG reasoning that incorporates asynchronous characteristics of event evolution into the model.
Outcome: The proposed model incorporates the asynchronous characteristics of event evolution for representation computation, thus improving prediction performance.
T-REx: A Large Scale Alignment of Natural Language with Knowledge Base Triples (L18-1)

Copied to clipboard

Challenge: Existing datasets that provide alignments between natural language and knowledge bases (KB) triples are limited in size, lack coverage and are of unreported quality.
Approach: They propose to build a large scale dataset of alignments between Wikipedia abstracts and Wikidata triples that is two orders of magnitude larger than the largest available alignments dataset.
Outcome: The proposed dataset is two orders of magnitude larger than the largest available dataset and covers 2.5 times more predicates.
NeuInfer: Knowledge Inference on N-ary Facts (2020.acl-main)

Copied to clipboard

Challenge: Existing studies on knowledge inference on binary facts have focused on finding out connotative valid facts.
Approach: They propose a neural network model, NeuInfer, for knowledge inference on n-ary facts.
Outcome: The proposed model can cope with the task to infer an unknown element in a whole fact, while ignoring the binary facts.
Hedwig: A Named Entity Linker (2020.lrec-1)

Copied to clipboard

Challenge: Named entity linking is the task of identifying mentions of named things in text . e.g., "Barack Obama" or "New York" are examples of named entities .
Approach: They propose an end-to-end named entity linker that uses BILSTM models for mention detection and a PageRank algorithm for entity linking.
Outcome: The proposed named entity linker performs better than the previous generation, and is trilingually better.
Program Transfer for Answering Complex Questions over Knowledge Bases (2022.acl-long)

Copied to clipboard

Challenge: Program induction for complex questions over knowledge bases relies on a large number of parallel question-program pairs for the given KB, but the gold program annotations are usually lacking, making learning difficult.
Approach: They propose an approach to leverage program annotations on rich KBs as external supervision signals to aid program induction for low-resourced KB.
Outcome: The proposed approach outperforms SOTA methods on ComplexWebQuestions and WebQuestionSP.
Select to Know: An Internal-External Knowledge Self-Selection Framework for Domain-Specific Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) perform well in general QA but often struggle in domain-specific scenarios.
Approach: They propose a framework that internalizes domain knowledge through internal-external knowledge self-selection and selective supervised fine-tuning.
Outcome: The proposed framework outperforms existing methods and matches domain-pretrained LLMs with significantly lower cost.
Knowledge-Enhanced Natural Language Inference Based on Knowledge Graphs (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to natural language inference rely on semantic knowledge, but background knowledge is limited to a few specific types.
Approach: They propose a Knowledge Graph-enhanced NLI model that leverages background knowledge stored in knowledge graphs to facilitate inference.
Outcome: The proposed model can leverage background knowledge stored in knowledge graphs to perform the task.
HyKGE: A Hypothesis Knowledge Graph Enhanced RAG Framework for Accurate and Reliable Medical LLMs Responses (2025.acl-long)

Copied to clipboard

Challenge: Recent approaches suffer from insufficient and repetitive knowledge retrieval, tedious and time-consuming query parsing, and monotonous knowledge utilization.
Approach: They propose a retrieval-augmented generation framework which leverages LLMs’ powerful reasoning capacity to compensate for the incompleteness of user queries.
Outcome: The proposed framework improves the accuracy and reliability of Large Language Models (LLMs) by combining the rich knowledge of LLMs with Hypothesis Outputs.
EDIN: An End-to-end Benchmark and Pipeline for Unknown Entity Discovery and Indexing (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work on Entity Linking assumes that the knowledge base is complete and all mentions can be linked.
Approach: They propose a temporally segmented Unknown Entity Discovery and Indexing (EDIN) benchmark where unknown entities have to be integrated into existing entity linking systems.
Outcome: The proposed system detects, clusters, and indexes mentions of unknown entities in context.
Ontology-Style Relation Annotation: A Case Study (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for Relation Extraction (RE) annotations use links between entities . a domain link connects the relation mention to the source entity while a range link connect the relation to the destination entity.
Approach: They propose an Ontology-Style Relation (OSR) annotation approach to find relation mentions in relation annotations.
Outcome: The proposed approach can be easily converted to Ontology RDF triples to populate an Ontologies.
KG-TRICK: Unifying Textual and Relational Information Completion of Knowledge for Multilingual Knowledge Graphs (2025.coling-main)

Copied to clipboard

Challenge: Existing studies have shown that combining information from KGs in different languages aids knowledge Graph Completion and Knowledge Graph Enhancement.
Approach: They propose a sequence-to-sequence framework that unifies tasks of textual and relational information completion for multilingual knowledge graphs.
Outcome: The proposed framework unifies tasks of KGC and KGE into a single framework.
From Zero to Hero: Human-In-The-Loop Entity Linking in Low Resource Domains (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to disambiguate entity mentions in a text depend on training data.
Approach: They propose a domain-agnostic approach to annotate entities using a KB-based approach.
Outcome: The proposed approach outperforms existing methods in a simulation on difficult texts.
EventOA: An Event Ontology Alignment Benchmark Based on FrameNet and Wikidata (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies on event ontologies focus on entity-based OA, and neglect event-based one . however, independent development of event ontoologies often results in heterogeneous representations that raise the need for establishing alignments between semantically related events.
Approach: They propose a multi-view event ontology alignment method that utilizes description information and neighbor information to obtain richer representations of the event ontoologies.
Outcome: The proposed method outperforms existing entity-based methods and can serve as a strong baseline for future research.
Improving Fine-grained Entity Typing with Entity Linking (D19-1)

Copied to clipboard

Challenge: Existing methods for fine-grained entity typing require a large tag set and knowledge of the context.
Approach: They propose a deep neural model that uses context and information from entity linking to improve fine-grained entity typing.
Outcome: The proposed model achieves 5% absolute strict accuracy improvement over the state of the art on two datasets.
Generating and Evaluating Plausible Explanations for Knowledge Graph Completion (2024.acl-long)

Copied to clipboard

Challenge: Existing XAI approaches focus on learning algorithmic explanations, but are not plausible for users.
Approach: They propose a path-based explanation method that meets human-centric explainability constraints and enhances plausibility.
Outcome: The proposed method meets human-centric explainability constraints and enhances plausibility.
From Text to Historical Ecological Knowledge: The Construction and Application of the Shan Jing Knowledge Base (2024.lrec-main)

Copied to clipboard

Challenge: Traditional Ecological Knowledge (TEK) is a shared cultural heritage and crucial instrument to tackle environmental challenges.
Approach: They propose to build a language resource based on Shanhai Jing (the classic of mountains and seas) written 2000 years ago and uses a stylized narrative and juxtaposition of knowledge from multiple domains to build the knowledge base.
Outcome: The proposed knowledge base contains 1432 systematically classified entities and 3294 relationships.
A Large Interlinked Knowledge Graph of the Italian Cultural Heritage (2022.lrec-1)

Copied to clipboard

Challenge: Existing efforts to create knowledge bases are limited to relatively small resources, such as entities from libraries, archeological sites and museums.
Approach: They propose to create a large knowledge graph linking Italian cultural heritage entities with concepts defined on well-known knowledge bases.
Outcome: The proposed graph shows that the Italian cultural heritage entities are interlinked with concepts defined on well-known knowledge bases.
mRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languages (2025.findings-acl)

Copied to clipboard

Challenge: Knowledge Graphs are structured multirelational graphs that store factual knowledge.
Approach: They introduce a Retrieval-Augmented Generation (mRAKL) based system to perform mKGC.
Outcome: The proposed approach improves over a no-context setting with an idealized retrieval system.
SpEL: Structured Prediction for Entity Linking (2023.emnlp-main)

Copied to clipboard

Challenge: Entity linking is a key component of structured data creation by linking spans of text to an ontology or knowledge source.
Approach: They propose to use structured prediction for entity linking to classify each input token as an entity and aggregate the token predictions.
Outcome: The proposed system outperforms the state-of-the-art on the commonly used AIDA benchmark dataset for entity linking to Wikipedia.
A Dataset for Hyper-Relational Extraction and a Cube-Filling Approach (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods do not consider qualifier attributes for each relation triplet, such as time, quantity or location.
Approach: They propose a hyper-relational extraction task to extract more specific facts from text using qualifiers.
Outcome: The proposed model outperforms baselines and reveal possible directions for future research.
Learning Collaborative Agents with Rule Guidance for Knowledge Graph Reasoning (2020.emnlp-main)

Copied to clipboard

Challenge: Walk-based models have shown their advantages in knowledge graph reasoning but are limited by their representations and generalizability.
Approach: They propose a walk-based model that leverages high-quality rules generated by symbolic-based methods to provide reward supervision for walk- based agents.
Outcome: Experiments on benchmark datasets show that RuleGuider improves the performance of walk-based models without losing interpretability.
Annotation Interoperability for the Post-ISOCat Era (2020.lrec-1)

Copied to clipboard

Challenge: Using ISOCat successor solutions, annotation standards have been developed since 2010 .
Approach: They describe ISOCat successor solutions and annotation standardization efforts since 2010 . they describe low-cost harmonization of post-ISOCat vocabularies by means of linked ontologies .
Outcome: The proposed ontologies are linked with the Ontologie of Linguistic Annotation and ISOCat, the GOLD ontology, the Typological Database Systems ontological and a large number of annotation schemes.
EnigmaToM: Improve LLMs’ Theory-of-Mind Reasoning Capabilities with Neural Knowledge Base of Entity States (2025.findings-acl)

Copied to clipboard

Challenge: Existing ToM reasoning methods rely excessively on off-the-shelf LLMs, reducing their efficiency and limiting their applicability to high-order ToM.
Approach: They propose a neuro-symbolic framework that integrates a Neural Knowledge Base of Entity States and knowledge injection to enhance ToM reasoning.
Outcome: The proposed framework improves ToM reasoning on ToMi, HiToM, and FANToM benchmarks.
BioFEG: Generate Latent Features for Biomedical Entity Linking (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to biomedical entity linking suffer from multiple types of errors due to the rarity of many biomedically relevant entities in real-world scenarios.
Approach: They propose a latent feature generation framework to generate latent semantic features for unseen entities to capture fine-grained coherence information of unseened entities.
Outcome: The proposed framework is superior to existing models on two benchmark datasets.
A Benchmark for Semi-Inductive Link Prediction in Knowledge Graphs (2023.findings-emnlp)

Copied to clipboard

Challenge: Semi-inductive link prediction (LP) is a task of predicting facts for new, previously unseen entities based on context information.
Approach: They propose to use Wikidata5M to evaluate semi-inductive link prediction (LP) in knowledge graphs.
Outcome: The proposed benchmark provides a test bed for further research into semi-inductive link prediction (LP) in knowledge graphs.
The Circumstantial Event Ontology (CEO) and ECB+/CEO: an Ontology and Corpus for Implicit Causal Relations between Events (L18-1)

Copied to clipboard

Challenge: a new ontology for calamity events models semantic circumstantial relations between event classes . a circumstancial relation makes clear "why" something happened, without necessarily predicting it.
Approach: They propose a circumstantial event ontology that models semantic circumstancial relations between event classes . they propose ECB+ annotated corpus for circumstantal relations and a meta model .
Outcome: The proposed model captures that the change yielded by one event explains to people the happening of the next event when observed.
Case-based Reasoning for Natural Language Queries over Knowledge Bases (2021.emnlp-main)

Copied to clipboard

Challenge: Using human-labeled examples, case-based reasoning can solve complex problems from scratch . case-Based reasoning is a paradigm that is used to solve complex problem .
Approach: They propose a neuro-symbolic CBR approach for question answering over large knowledge bases.
Outcome: The proposed approach outperforms the current state of the art on a CWQ dataset by 11% on accuracy.
Argument-Aware Approach To Event Linking (2024.findings-acl)

Copied to clipboard

Challenge: Prior research in event linking has mainly borrowed methods from entity linking, overlooking distinct features of events.
Approach: They propose an argument-aware method to improve event linking models by augmenting input text with tagged event argument information.
Outcome: The proposed method improves in-KB and out-of-KB queries and training examples.
iQUEST: An Iterative Question-Guided Framework for Knowledge Base Question Answering (2025.acl-long)

Copied to clipboard

Challenge: Large language models suffer from factual inaccuracies in knowledge-intensive domains.
Approach: They propose a question-guided KBQA framework that iteratively decomposes complex queries into simpler sub-questions and integrates a Graph Neural Network (GNN) to look ahead and incorporate 2-hop neighbor information at each reasoning step.
Outcome: The proposed framework improves on four benchmark datasets and four LLMs.
DynamicER: Resolving Emerging Mentions to Dynamic Entities for RAG (2024.emnlp-main)

Copied to clipboard

Challenge: Existing entity linking models struggle to link new expressions to entities in the dynamic nature of human language.
Approach: They propose a task to resolve emerging mentions to dynamic entities and a benchmark to evaluate their model's adaptability to new expressions.
Outcome: The proposed method outperforms baselines on QA task with resolved mentions and improves retrieval-augmented generation performance.
ReLiK: Retrieve and LinK, Fast and Accurate Entity Linking and Relation Extraction on an Academic Budget (2024.findings-acl)

Copied to clipboard

Challenge: Entity Linking and Relation Extraction (EL) are fundamental tasks in Natural Language Processing.
Approach: They propose a Retriever-Reader architecture for Entity Linking and Relation Extraction . they propose an input representation that incorporates the candidate entities alongside the text .
Outcome: The proposed architecture achieves state-of-the-art in in- and out-of domain benchmarks while using academic budget training and with 40x inference speed compared to competitors.
Knowledge-aware Attention Network for Medication Effectiveness Prediction (2024.lrec-main)

Copied to clipboard

Challenge: Existing effectiveness prediction methods focus on one specific medicine, one specific disease, or one specific lab test, making it hard to extend to general medicines and diseases in hospital/ICU scenarios.
Approach: They propose to use knowledge enhanced module to incorporate external knowledge about medications and a medical feature learning module to determine the interaction between diagnosis and medications.
Outcome: The proposed model outperforms state-of-the-art methods on a public dataset showing that it significantly outperformed existing models.
Knowledge Graphs for Real-World Rumour Verification (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in automated rumour verification have limited results in real-world scenarios.
Approach: They propose to use Twitter responses to construct knowledge graphs based on the PHEME dataset to identify discrepancies between the evidence retrieved and PHE ME’s labels.
Outcome: The proposed model outperforms the state-of-the-art on PHEME and has superior generisability when evaluated on a temporally distant rumour verification dataset.
Editing OntoLex-Lemon in VocBench 3 (2020.lrec-1)

Copied to clipboard

Challenge: OntoLex-Lemon is a collection of RDF vocabularies for specifying the verbalization of ontologies in natural language.
Approach: They propose to extend existing RDF editor to OntoLex-Lemon to provide more direct editing . they propose to use a model that allows for the verbalization of ontologies in natural language .
Outcome: The proposed editor improves the ontology-lexicon interface and improves its flexibility.
A Comprehensive Evaluation of Biomedical Entity Linking Models (2023.emnlp-main)

Copied to clipboard

Challenge: Current methods struggle to correctly link genes and proteins and often have difficulty incorporating context into linking decisions.
Approach: They evaluate nine recent state-of-the-art biomedical entity linking models under a unified framework.
Outcome: The proposed models are compared along axes of accuracy, speed, ease of use, generalization, adaptability and adaptability to new ontologies and datasets.
DAVIS: Planning Agent with Knowledge Graph-Powered Inner Monologue (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to designing a generalist scientific agent fail to address multifaceted requirements of scientific tasks.
Approach: DAVIS is a generalist scientific agent capable of performing tasks in laboratory settings to assist researchers.
Outcome: DAVIS performs better on ScienceWorld benchmarks compared to previous approaches on 8 out of 9 elementary science subjects.
Distance-Based Propagation for Efficient Knowledge Graph Reasoning (2023.emnlp-main)

Copied to clipboard

Challenge: Knowledge graph completion (KGC) aims to predict unseen edges in knowledge graphs (KGs) . a few recent attempts to address this problem sacrifice the performance to gain efficiency.
Approach: They propose a method that aggregates path information to solve this problem by aggregating paths in a fixed window for each source-target pair.
Outcome: The proposed method can cut down on the number of propagated messages by 90% while achieving competitive performance on multiple KG datasets.
Detecting Spoilers in Movie Reviews with External Movie Knowledge and User Networks (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models focus on the textual content of the review, while spoiler detection requires putting the review into the context of facts and knowledge regarding movies.
Approach: They propose a network-based spoiler detection model that takes into account external knowledge about movies and user activities on movie review platforms.
Outcome: The proposed model takes into account external knowledge about movies and user activities on movie review platforms while incorporating user networks.
mReFinED: An Efficient End-to-End Multilingual Entity Linking System (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing work assumed that entity mentions were given and skipped the entity mention detection step due to a lack of high-quality multilingual training corpora.
Approach: They propose a bootstrapping mention detection framework that enhances the quality of training corpora.
Outcome: The proposed framework outperforms existing work in the end-to-end MEL task while being 44 times faster.
Modelling and Linking an Old Latin-Portuguese Dictionary to the LiLa Knowledge Base (2024.lrec-main)

Copied to clipboard

Challenge: lexical and lexicographic information of Antonio Velez's bilingual Latin-Portuguese dictionary was modelled using the Lexicon Model for Ontologies and its lexicog module.
Approach: This paper describes steps undertaken to include data from Antonio Velez’s bilingual Latin-Portuguese dictionary into the LiLa Knowledge Base of interoperable linguistic resources for Latin.
Outcome: The proposed model includes lexical and lexicographic information from the source dictionary with those of the LiLa collection of Latin lemmas.
DLTKG: Denoising Logic-based Temporal Knowledge Graph Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: Current approaches to temporal knowledge representation face limited generalization to unseen facts and insufficient interpretability of reasoning processes.
Approach: They propose a framework that uses a denoising diffusion process to complete reasoning tasks . they propose introducing a noise source and historical conditionguiding mechanism to improve interpretability .
Outcome: The proposed framework outperforms state-of-the-art methods on three benchmark datasets.
A Multi-Expert Structural-Semantic Hybrid Framework for Unveiling Historical Patterns in Temporal Knowledge Graphs (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods focus on graph structure learning or semantic reasoning, lacking the capability to capture the inherent differences between historical and non-historical events.
Approach: They propose a temporal knowledge graph reasoning framework that integrates both structural and semantic information to guide the reasoning process for different events.
Outcome: The proposed framework integrates structural and semantic information to predict future events . it can provide evidence for many downstream tasks, including situation analysis and political decision making .
MedCPI: A Construct–Personalize–Integrate Framework for KG-enhanced Clinical Prediction (2026.findings-acl)

Copied to clipboard

Challenge: Existing KG-enhanced approaches to clinical prediction are limited . existing approaches to personalize and integrate knowledge are weakly controlled .
Approach: They propose a framework to integrate medical knowledge graphs into EHRs to support KG-enhanced clinical prediction.
Outcome: The proposed framework improves on MIMIC-III and MIMIC IV tasks.
Selective Temporal Knowledge Graph Reasoning (2024.lrec-main)

Copied to clipboard

Challenge: Existing models cannot abstain from uncertain predictions, which will bring risks in real-world applications.
Approach: They propose to abstain from uncertain future facts by using a confidence estimator . they take both the certainty of the current prediction and the accuracy of historical predictions into account .
Outcome: The proposed abstention mechanism helps existing models make selective predictions instead of indiscriminate ones.
Sequential and Repetitive Pattern Learning for Temporal Knowledge Graph Reasoning (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to learn temporal evolutional representations of entities are hard to capture the complex temporal patterns such as sequential and repetitive.
Approach: They propose a Sequential and Repetitive Pattern Learning method that captures both sequential and repetitive patterns.
Outcome: The proposed method outperforms state-of-the-art methods on four representative benchmarks on GDELT dataset, where performance improvement of MRR reaches up to 18.84%.
A Survey of Link Prediction in N-ary Knowledge Graphs (2025.emnlp-main)

Copied to clipboard

Challenge: N-ary Knowledge Graphs (NKGs) capture n-ary facts containing more than two entities.
Approach: They present the first comprehensive survey of link prediction in NKGs . they provide an overview of the field and analyze their performance and application scenarios .
Outcome: The proposed methods provide an overview of the field and analyze performance and application scenarios.
NeuRAG: End-to-End Neural Knowledge Augmentation via Hyper-Neurons (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to grounding large language models in external knowledge are constrained by a decoupled architecture: retrieval and reasoning operate as separate stages, with retrieved text merely prepended as passive context.
Approach: They propose an end-to-end Neuralized RAG framework that unifies knowledge retrieval and fusion through Hyper-Neurons.
Outcome: Extensive experiments across multiple datasets and LLMs demonstrate NeuRAG’s strong and consistent performance as a promising novel RAG paradigm.
EvoMemKG: An Evolvable Memory Agent for Multi-hop Knowledge Graph Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for integrating knowledge graphs with large language models lack continuous learning capabilities.
Approach: They propose an agent framework with a dynamic, evolvable memory mechanism specifically designed for KG reasoning.
Outcome: EvoMemKG achieves state-of-the-art performance without training or tools . it achieves improvements of up to 20% over baseline on multi-hop queries .
RAED: Retrieval-Augmented Entity Description Generation for Emerging Entity Linking and Disambiguation (2025.emnlp-main)

Copied to clipboard

Challenge: Entity Linking and Entity Disambiguation systems assume static knowledge bases are incomplete and up-to-date, rendering them incapable of handling entities not yet included in the knowledge base.
Approach: They propose a model that retrieves external knowledge to improve factual grounding in entity descriptions.
Outcome: The proposed model outperforms systems that require fixed knowledge sets on Entity Disambiguation and Wikipedia to improve factual grounding in entity descriptions.
Chain-of-Relations: Faithful and Efficient LLM Reasoning over Knowledge Graphs via Relation-Centric Exploration (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods adopt entity-centric exploration that incrementally constructs reasoning paths by selecting and connecting intermediate entities.
Approach: They propose to use relation-centric exploration to construct reasoning paths by selecting and connecting intermediate entities and to reduce the dependence on entity completeness.
Outcome: The proposed method outperforms baselines on three benchmark datasets in both F1 score and KG-grounded Rate.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations